Index-Translate-2B-FP8

Official FP8 (W8A8) quantization of IndexTeam/Index-Translate-2B, part of the Index-Translate multilingual translation model family (150 languages, terminology/format-constrained translation, controlled dubbing translation, long-document translation).

Quantization

  • Scheme: FP8_DYNAMIC (FP8 E4M3 weights, per-token dynamic FP8 activations), produced with llm-compressor (quantization_scheme recorded in recipe.yaml).
  • All Linear layers of the language model are quantized; the vision tower, multi-modal projector, lm_head, and embeddings are kept in BF16. MTP (multi-token prediction) weights are preserved.
  • Format: compressed-tensors safetensors — load directly with vLLM (quantization="compressed-tensors") or transformers.

Consistency validation

Measured on an NVIDIA A100 against the original BF16 checkpoint (greedy decoding, official translation prompt):

Metric BF16 FP8 Delta
Perplexity (fixed corpus) 3.9104 3.9403 +0.76%
zh→en generation identical — — yes
en→zh generation identical — — no (semantically equivalent)

Usage

vllm serve IndexTeam/Index-Translate-2B-FP8 --host 127.0.0.1 --port 8000 --max-model-len 4096

Translation prompt format (greedy decoding, temperature=0 recommended; chat template with enable_thinking: false):

请将以下文本翻译为{target-language},直接输出翻译结果,不要进行任何解释。

{source-text}

Prompting & constrained translation (instTrans)

Beyond plain translation, the models follow the instTrans constrained-translation format. The official client wraps requests into the canonical structure 【源文】<text> + numbered 1. 【硬性要求】<hard constraints> + 2. 【注意】<soft constraints> + suffix instructions:

  • Hard constraints (binary, must hold): strict terminology glossary enforcement (e.g. 碳纤维:carbon fiber, 抗裂缝:crack resistance), and format/structure preservation for JSON/CSV/code/placeholders.
  • Soft constraints (graded): tone & style adaptation (e.g. formal business-email register), domain/word-sense disambiguation (e.g. plant → 工厂 in an industrial context), cross-sentence consistency, LaTeX preservation.
  • Syllable-controlled translation (dubbing): use the Index-Homura checkpoints (IndexTeam/Index-Homura-2B/9B), which strictly respect a target syllable budget and can be combined with glossaries.

Full prompt reference: github.com/bilibili/Index-Translate (Instruction Following section, docs/prompts.md, inference/llm/cases/).

See the base model card for the full instTrans constrained-translation format and serving presets. GGUF builds for local inference are published in Index-Translate-2B-GGUF.

Quantized and published by the Index team, 2026-10-03.

Downloads last month
18
Safetensors
Model size
2B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IndexTeam/Index-Translate-2B-FP8

Quantized
(16)
this model

Collection including IndexTeam/Index-Translate-2B-FP8

Paper for IndexTeam/Index-Translate-2B-FP8