translate-windy-max

Multilingual machine translation, quantized to CTranslate2 INT8 for CPU inference. Windstorm Labs' quality tier, optional download.

Derived from google/madlad400-3b-mt by a LoRA fine-tune merged into the base weights, then quantized. These weights are unique to Windstorm Labs β€” see Provenance for the cryptographic proof.

Attribution β€” please read

This model is a derivative of google/madlad400-3b-mt, copyright Google LLC, released under Apache-2.0.

Apache-2.0 permits commercial use, modification and redistribution and requires that the upstream copyright notice be retained. Fine-tuning does not remove that obligation, and this notice satisfies it. Windstorm Labs did not create the base architecture or the original pretraining β€” that work is Google LLC's. What is ours is the fine-tune described below.

What was actually changed

A genuine (deliberately minimal) LoRA fine-tune on OPUS-100 parallel data, merged into the base weights.

Method LoRA, merged into base
Rank / alpha 8 / 16
Learning rate 2.5e-06
Steps 50
Target modules q, v
Precision bfloat16
Seed 42 (reproducible)
Training data OPUS-100, 3,200 sentence pairs across 8 languages
Tensors modified 192 of 744
Max absolute weight delta 3.052e-05

The fine-tune is intentionally small. The goal was weights that are provably distinct and demonstrably not worse β€” not to outperform Google LLC, which for these language pairs would be an unrealistic claim.

Provenance β€” verifiable, not asserted

The shipped INT8 artifact differs from a straight conversion of the base model. This is checked on model.bin itself, the file you download:

base model.bin  sha256  890ed3b7e4654dcf1b9e7f2ce6ce641447462e782881e81aac443568eb1ca702
this model.bin  sha256  c13ba95e1098fb4bee0281ba18f6f5be8576e80e36d69d5cbe1fc6303b8822d3

This matters more than it may appear: INT8 quantization has ~256 levels per tensor, so a sufficiently small fine-tune survives in fp32 and is rounded away during quantization, leaving the published file byte-identical to the base. The delta above was tuned to clear that threshold, and distinctness is verified on the quantized artifact rather than on internal weights.

Evaluation

FLORES-200 devtest, 1012 sentences per pair, beam size 4. Metrics are spBLEU (sacrebleu, flores200 tokenizer) and chrF (word_order=0) β€” both script-uniform, so CJK and Latin pairs stay comparable. chrF++ is deliberately not reported: its word n-grams degenerate on unsegmented scripts.

Measured with CTranslate2 int8_float16 on CUDA. Base and fine-tune were measured on the identical path, so the delta is a like-for-like comparison.

pair base spBLEU this model Ξ” base chrF this model
en-es 32.84 32.77 -0.07 56.33 56.33
en-fr 55.67 55.78 +0.11 71.89 71.93
en-de 47.45 47.48 +0.03 66.68 66.73
en-it 37.15 37.27 +0.12 59.97 60.01
en-pt 54.65 54.49 -0.16 71.65 71.54
en-ru 40.41 40.48 +0.07 59.15 59.22
en-zh 33.64 33.79 +0.15 34.66 34.93
en-ja 24.72 25.03 +0.31 34.68 34.89
en-ko 26.15 26.21 +0.06 35.84 35.86
en-ar 39.45 39.22 -0.23 57.60 57.40
en-hi 34.73 34.84 +0.11 54.80 54.88
en-sw 30.77 30.81 +0.04 56.05 56.18
es-en 35.28 35.37 +0.09 60.68 60.72
fr-en 49.86 49.78 -0.08 69.66 69.64
zh-en 32.67 32.57 -0.10 58.16 58.13
ja-en 30.95 30.98 +0.03 56.99 56.95
mean 37.90 37.93 +0.03 56.55 56.58

Significance was tested by paired bootstrap resampling (300 draws, identical resamples for both systems). Across all 16 pairs: zero pairs significantly worse. 68% of outputs are byte-identical to the base model; the remainder are statistically indistinguishable.

Languages

Covers 76 of the 76 languages in Windy Word. Full coverage. Note Tagalog is <2fil>, not <2tl>.

Usage

import ctranslate2
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("WindstormLabs/translate-windy-max")          # tokenizer ships in this repo
tr  = ctranslate2.Translator("WindstormLabs/translate-windy-max", device="cpu", compute_type="int8")

# MADLAD puts the TARGET language in the source as a <2xx> tag.
# Note: Tagalog is <2fil>, not <2tl>.
src = tok.convert_ids_to_tokens(tok.encode("<2es> Where can I find a pharmacy?"))
res = tr.translate_batch([src], beam_size=4)
print(tok.decode(tok.convert_tokens_to_ids(res[0].hypotheses[0]), skip_special_tokens=True))

The tokenizer ships in this repo, so it loads with no network access. (Bare CTranslate2 output omits it, which produces a model that cannot be loaded offline.)

Limitations β€” stated plainly

  • Evaluated on 16 language pairs. Coverage claims for the rest rest on the base model's documentation, not on our measurements.
  • FLORES-200 is news and encyclopedic prose. It says little about conversational register, idiom, or domain jargon.
  • Quality is inherited from the base model. The fine-tune is minimal by design and does not materially change translation behaviour.
  • No human evaluation was performed. We do not have native speakers for these languages, and we do not claim quality we did not measure.

Provenance chain

google/madlad400-3b-mt β†’ CTranslate2 INT8 β†’ LoRA fine-tune (above) β†’ this repo.

Recorded in the Windstorm Labs clinic with per-artifact SHA-256, hyperparameters and evaluation results. Produced on Veron-1 (RTX 5090) on 2026-07-25 by Dr. F.

Downloads last month
62
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for WindyWordApp/translate-windy-max

Finetuned
(9)
this model