Instructions to use IndexTeam/Index-Translate-9B-FP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IndexTeam/Index-Translate-9B-FP4 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="IndexTeam/Index-Translate-9B-FP4")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("IndexTeam/Index-Translate-9B-FP4") model = AutoModelForMultimodalLM.from_pretrained("IndexTeam/Index-Translate-9B-FP4", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Index-Translate-9B-FP4
Official NVFP4 (W4A4) quantization of IndexTeam/Index-Translate-9B, part of the Index-Translate multilingual translation model family (150 languages, terminology/format-constrained translation, controlled dubbing translation, long-document translation).
- Technical report: Index-Translate: A Multilingual Translation Model Family
- Code: github.com/bilibili/Index-Translate
Quantization
- Scheme:
NVFP4(W4A4: 4-bit floating-point weights with per-group-16 scales, 4-bit floating-point activations with calibrated per-tensor global scales), produced with llm-compressor (quantization_schemerecorded inrecipe.yaml). Calibrated on a small bilingual translation corpus. - All
Linearlayers of the language model are quantized;lm_head, embeddings, MoE router gates and shared-expert gates are kept in BF16. MTP (multi-token prediction) weights are preserved. - Format: compressed-tensors
nvfp4-pack-quantizedsafetensors - load directly with vLLM (quantization="compressed-tensors") or transformers.
Consistency validation
Measured on an NVIDIA A100 (weight-dequantized execution) against the original BF16 checkpoint (greedy decoding, official translation prompt):
| Metric | BF16 | FP4 | Delta |
|---|---|---|---|
| Perplexity (fixed corpus) | 2.5370 | 2.5949 | +2.28% |
| zh->en generation identical | - | - | yes |
| en->zh generation identical | - | - | no (semantically equivalent) |
Usage
vllm serve IndexTeam/Index-Translate-9B-FP4 --host 127.0.0.1 --port 8000 --max-model-len 4096
Hardware note: full NVFP4 (W4A4) acceleration requires an NVIDIA Blackwell GPU (SM100+, e.g. B200 / RTX 50 series). On older GPUs (Hopper / Ampere) vLLM loads the checkpoint with weight-only dequantization - memory is still reduced, but there is no FP4 compute speedup. For non-Blackwell serving we recommend the FP8 build.
Translation prompt format (greedy decoding, temperature=0 recommended; chat template with enable_thinking: false):
请将以下文本翻译为{target-language},直接输出翻译结果,不要进行任何解释。
{source-text}
Prompting & constrained translation (instTrans)
Beyond plain translation, the models follow the instTrans constrained-translation format. The official client wraps requests into the canonical structure 【源文】<text> + numbered 1. 【硬性要求】<hard constraints> + 2. 【注意】<soft constraints> + suffix instructions:
- Hard constraints (binary, must hold): strict terminology glossary enforcement (e.g.
碳纤维:carbon fiber, 抗裂缝:crack resistance), and format/structure preservation for JSON/CSV/code/placeholders. - Soft constraints (graded): tone & style adaptation (e.g. formal business-email register), domain/word-sense disambiguation (e.g. plant -> 工厂 in an industrial context), cross-sentence consistency, LaTeX preservation.
- Syllable-controlled translation (dubbing): the Index-Homura checkpoints (IndexTeam/Index-Homura-2B, IndexTeam/Index-Homura-9B) strictly respect a target syllable budget and can be combined with glossaries.
Full prompt reference: github.com/bilibili/Index-Translate (Instruction Following section, docs/prompts.md, inference/llm/cases/).
See the base model card for the full instTrans constrained-translation format and serving presets. GGUF builds for local inference are published in Index-Translate-9B-GGUF, and an FP8 build for Hopper/Ampere serving in Index-Translate-9B-FP8.
Quantized and published by the Index team, 2026-10-04.
- Downloads last month
- 23
Model tree for IndexTeam/Index-Translate-9B-FP4
Base model
IndexTeam/Index-Translate-9B