Index-Translate-9B-FP4

Official NVFP4 (W4A4) quantization of IndexTeam/Index-Translate-9B, part of the Index-Translate multilingual translation model family (150 languages, terminology/format-constrained translation, controlled dubbing translation, long-document translation).

Quantization

  • Scheme: NVFP4 (W4A4: 4-bit floating-point weights with per-group-16 scales, 4-bit floating-point activations with calibrated per-tensor global scales), produced with llm-compressor (quantization_scheme recorded in recipe.yaml). Calibrated on a small bilingual translation corpus.
  • All Linear layers of the language model are quantized; lm_head, embeddings, MoE router gates and shared-expert gates are kept in BF16. MTP (multi-token prediction) weights are preserved.
  • Format: compressed-tensors nvfp4-pack-quantized safetensors - load directly with vLLM (quantization="compressed-tensors") or transformers.

Consistency validation

Measured on an NVIDIA A100 (weight-dequantized execution) against the original BF16 checkpoint (greedy decoding, official translation prompt):

Metric BF16 FP4 Delta
Perplexity (fixed corpus) 2.5370 2.5949 +2.28%
zh->en generation identical - - yes
en->zh generation identical - - no (semantically equivalent)

Usage

vllm serve IndexTeam/Index-Translate-9B-FP4 --host 127.0.0.1 --port 8000 --max-model-len 4096

Hardware note: full NVFP4 (W4A4) acceleration requires an NVIDIA Blackwell GPU (SM100+, e.g. B200 / RTX 50 series). On older GPUs (Hopper / Ampere) vLLM loads the checkpoint with weight-only dequantization - memory is still reduced, but there is no FP4 compute speedup. For non-Blackwell serving we recommend the FP8 build.

Translation prompt format (greedy decoding, temperature=0 recommended; chat template with enable_thinking: false):

请将以下文本翻译为{target-language},直接输出翻译结果,不要进行任何解释。

{source-text}

Prompting & constrained translation (instTrans)

Beyond plain translation, the models follow the instTrans constrained-translation format. The official client wraps requests into the canonical structure 【源文】<text> + numbered 1. 【硬性要求】<hard constraints> + 2. 【注意】<soft constraints> + suffix instructions:

  • Hard constraints (binary, must hold): strict terminology glossary enforcement (e.g. 碳纤维:carbon fiber, 抗裂缝:crack resistance), and format/structure preservation for JSON/CSV/code/placeholders.
  • Soft constraints (graded): tone & style adaptation (e.g. formal business-email register), domain/word-sense disambiguation (e.g. plant -> 工厂 in an industrial context), cross-sentence consistency, LaTeX preservation.
  • Syllable-controlled translation (dubbing): the Index-Homura checkpoints (IndexTeam/Index-Homura-2B, IndexTeam/Index-Homura-9B) strictly respect a target syllable budget and can be combined with glossaries.

Full prompt reference: github.com/bilibili/Index-Translate (Instruction Following section, docs/prompts.md, inference/llm/cases/).

See the base model card for the full instTrans constrained-translation format and serving presets. GGUF builds for local inference are published in Index-Translate-9B-GGUF, and an FP8 build for Hopper/Ampere serving in Index-Translate-9B-FP8.

Quantized and published by the Index team, 2026-10-04.

Downloads last month
23
Safetensors
Model size
9B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IndexTeam/Index-Translate-9B-FP4

Quantized
(12)
this model

Collection including IndexTeam/Index-Translate-9B-FP4

Paper for IndexTeam/Index-Translate-9B-FP4