JMangaTranslator-UltraFast-Exp

English | 中文

Experimental. For actual translation, use JMangaTranslator-Fast instead.

JMangaTranslator-UltraFast-Exp translates Japanese manga speech bubbles into Simplified Chinese, one bubble at a time. It is a 229M-parameter non-autoregressive model (a Directed Acyclic Transformer, DAT): it produces the whole translation in one forward pass. On an RTX 3060 it takes 4.5 ms per bubble, about twice as fast as JMangaTranslator-Fast v1. Its quality, however, is clearly lower than v1's.

We publish it as a negative result: the speed is real, but the quality gap did not close. The model, the measurements and the open problem below may be useful to anyone working on non-autoregressive translation.

Architecture, training, evaluation and speed details are in the technical report.

What we learned

  • Faster, but worse. On 3,817 Manga109-s bubbles, v1 beats this model by 0.029 COMET on ground-truth text and 0.038 on manga-ocr output. Every paired 95% bootstrap interval excludes zero. In exchange, this model saves about 4 ms per bubble on an RTX 3060 (4.5 ms vs 8.6 ms). v1 is already fast enough to translate each bubble as soon as it is recognized, so we consider the trade not worth it.
  • Manga fine-tuning helped only a little. The base model was trained from scratch on light-novel data. Fine-tuning it on 250k manga bubbles raised COMET by 0.004 on both input conditions, a significant gain, but small next to the gap to v1.
  • Expensive to train. The base model took 30.1 hours on an RTX 5090. JMangaTranslator-Fast v1, which translates better, took about 15 hours on an RTX 4090 in total. A DAT expands every training example into a graph four times as long as the source, so each training step does much more work than in an ordinary encoder–decoder model.
  • The better translation is often already in the graph. A DAT predicts a graph of candidate tokens and then picks one path through it. On two other test sets (Manga200 and Luna1k), we took the 12 beam-search paths from the same forward pass and chose the best one with the reference translation. This oracle choice was 0.043 COMET better than the default path on both sets. Without the reference, the best reranker we tried recovered only +0.020 and +0.0055. These numbers are for the base model before fine-tuning. Picking a better path without a reference, at low cost, is the open problem.

Quality

COMET-22 on 3,817 Manga109-s bubbles, the same set and protocol as the JMangaTranslator-Fast v1 results. The set keeps every bubble on which an OCR model made a mistake, so it is harder than average. "Ground truth" is the annotated text; "manga-ocr" is the same bubbles as read by manga-ocr. Reference translations were produced by Claude.

System Parameters Ground truth manga-ocr
JMangaTranslator-Fast v1 368M 0.8879 0.8315
GalTransl-v4-4B 4B 0.8617 0.8022
JMangaTranslator-UltraFast-Exp 229M 0.8588 0.7932
(the same model before manga fine-tuning) 229M 0.8543 0.7890
NanoSakura-2.2-0.2B 0.2B 0.8456 0.7827
Sakura-1.5B-Qwen2.5-v1.0 1.5B 0.8392 0.7787
Hy-MT2-1.8B-JP-Manga-Finetune-zh-Hans-v1 1.8B 0.8349 0.7850
Hy-MT2-1.8B 1.8B 0.8299 0.7818
opus-mt-ja-zh 77M 0.6849 0.6532
M2M100-418M 418M 0.6582 0.6278
NLLB-200-distilled-600M 600M 0.6122 0.5904

Speed

RTX 3060 12 GB, batch size 1, 500 Manga109-s bubbles, from tokenization to the decoded string. Load time includes building the CUDA graphs.

Model Backend p50 p90 Load time
UltraFast-Exp CUDA Graphs, fp16 (default) 4.5 ms 5.8 ms 7.9 s
UltraFast-Exp CUDA Graphs + torch.compile, fp16 2.8 ms 3.9 ms several minutes (260 s in our test)
UltraFast-Exp PyTorch eager, fp32 37.4 ms 38.6 ms 6.0 s
JMangaTranslator-Fast v1 CUDA Graphs, fp16 8.6 ms 12.8 ms 3.4 s

The compiled backend fuses many small GPU kernels and saves another 1.7 ms per bubble, but it must compile for several minutes at every first start. It is therefore off by default.

Usage

The weights and code are on Hugging Face and ModelScope. GitHub holds the code only.

hf download muscgab/JMangaTranslator-UltraFast-Exp --local-dir JMangaTranslator-UltraFast-Exp
# or: modelscope download --model muscgab/JMangaTranslator-UltraFast-Exp --local_dir JMangaTranslator-UltraFast-Exp
cd JMangaTranslator-UltraFast-Exp
pip install -r requirements.txt
python translate.py "堪忍袋の緒が切れた!"
python translate.py < bubbles.txt > translations.txt
python translate.py --compile < bubbles.txt > translations.txt   # NVIDIA only; compiles for several minutes first
from jmt_ultrafast import load
tr = load("path/to/JMangaTranslator-UltraFast-Exp")   # add compile=True for the compiled backend
print(tr.translate("堪忍袋の緒が切れた!"))   # 忍耐断了!  (v1: 忍无可忍了!)

Give one bubble per line, with line breaks inside a bubble removed. On an NVIDIA GPU the CUDA Graphs backend is used. Elsewhere the model runs in PyTorch eager mode, which is much slower.

The CPU path of the packaged code was checked bubble by bubble against the measured outputs. The CUDA Graphs and compiled paths have not yet been re-run in this packaged form; the speed numbers above come from the research version of the same code.

Limitations

  • Quality. Lower than v1 on both input conditions (see above).
  • Typical errors. On the 3,817 ground-truth bubbles, 18 outputs still contain Japanese kana and 7 contain the replacement character U+FFFD. v1 has none of either.
  • No context. Each bubble is translated on its own, so names, pronouns and tone may differ between bubbles.
  • Machine-made training targets. No human translations were used in training.

Acknowledgments

License

  • Model weights: CC BY-NC-SA 4.0 (LICENSE-weights.md). Attribution is required, commercial use is not permitted, and adapted models must be shared under the same license.
  • Code: MIT (LICENSE).

Training used OpenSakura data (license "other", intended for research and model development) and the author's private manga text. Manga109-s was used for evaluation only. No weights or outputs of the other systems in the comparison are included.

Downloads last month
12
Safetensors
Model size
0.2B params
Tensor type
I64
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for muscgab/JMangaTranslator-UltraFast-Exp