Instructions to use IndexTeam/Index-Translate-9B-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IndexTeam/Index-Translate-9B-FP8 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("translation", model="IndexTeam/Index-Translate-9B-FP8")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("IndexTeam/Index-Translate-9B-FP8") model = AutoModelForMultimodalLM.from_pretrained("IndexTeam/Index-Translate-9B-FP8", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Index-Translate-9B-FP8
Official FP8 (W8A8) quantization of IndexTeam/Index-Translate-9B, part of the Index-Translate multilingual translation model family (150 languages, terminology/format-constrained translation, controlled dubbing translation, long-document translation).
- Technical report: Index-Translate: A Multilingual Translation Model Family
- Code: github.com/bilibili/Index-Translate
Quantization
- Scheme:
FP8_DYNAMIC(FP8 E4M3 weights, per-token dynamic FP8 activations), produced with llm-compressor (quantization_schemerecorded inrecipe.yaml). - All
Linearlayers of the language model are quantized; the vision tower, multi-modal projector,lm_head, and embeddings are kept in BF16. MTP (multi-token prediction) weights are preserved. - Format: compressed-tensors safetensors — load directly with vLLM (
quantization="compressed-tensors") or transformers.
Consistency validation
Measured on an NVIDIA A100 against the original BF16 checkpoint (greedy decoding, official translation prompt):
| Metric | BF16 | FP8 | Delta |
|---|---|---|---|
| Perplexity (fixed corpus) | 2.7722 | 2.7838 | +0.42% |
| zh→en generation identical | — | — | yes |
| en→zh generation identical | — | — | yes |
Usage
vllm serve IndexTeam/Index-Translate-9B-FP8 --host 127.0.0.1 --port 8000 --max-model-len 4096
Translation prompt format (greedy decoding, temperature=0 recommended; chat template with enable_thinking: false):
请将以下文本翻译为{target-language},直接输出翻译结果,不要进行任何解释。
{source-text}
Prompting & constrained translation (instTrans)
Beyond plain translation, the models follow the instTrans constrained-translation format. The official client wraps requests into the canonical structure 【源文】<text> + numbered 1. 【硬性要求】<hard constraints> + 2. 【注意】<soft constraints> + suffix instructions:
- Hard constraints (binary, must hold): strict terminology glossary enforcement (e.g.
碳纤维:carbon fiber, 抗裂缝:crack resistance), and format/structure preservation for JSON/CSV/code/placeholders. - Soft constraints (graded): tone & style adaptation (e.g. formal business-email register), domain/word-sense disambiguation (e.g. plant → 工厂 in an industrial context), cross-sentence consistency, LaTeX preservation.
- Syllable-controlled translation (dubbing): use the Index-Homura checkpoints (IndexTeam/Index-Homura-2B/9B), which strictly respect a target syllable budget and can be combined with glossaries.
Full prompt reference: github.com/bilibili/Index-Translate (Instruction Following section, docs/prompts.md, inference/llm/cases/).
See the base model card for the full instTrans constrained-translation format and serving presets. GGUF builds for local inference are published in Index-Translate-9B-GGUF.
Quantized and published by the Index team, 2026-10-03.
- Downloads last month
- 28
Model tree for IndexTeam/Index-Translate-9B-FP8
Base model
IndexTeam/Index-Translate-9B