Instructions to use TigreGotico/opus-mt-tr-az-onnx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TigreGotico/opus-mt-tr-az-onnx with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("translation", model="TigreGotico/opus-mt-tr-az-onnx")# Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("TigreGotico/opus-mt-tr-az-onnx") model = AutoModelForSeq2SeqLM.from_pretrained("TigreGotico/opus-mt-tr-az-onnx", device_map="auto") - Notebooks
- Google Colab
- Kaggle
opus-mt-tr-az-onnx
ONNX export (fp32 + dynamic int8) of Helsinki-NLP/opus-mt-tr-az,
a Marian (tr -> az) translation model from the Helsinki-NLP OPUS-MT project.
Completes a Turkic mini-cluster alongside the existing en-tr/tr-en and the
en-az/az-en pair converted alongside it.
License: apache-2.0, inherited unchanged from the base model.
Contents
| Files | What it is |
|---|---|
encoder_model.onnx, decoder_model.onnx, decoder_with_past_model.onnx |
ONNX, float32 |
int8/ |
same three graphs, dynamic int8 (QInt8, MatMul only, /lm_head/MatMul excluded) |
source.spm, target.spm, vocab.json |
MarianTokenizer over the original sentencepiece models |
fp32 size: 325 MB (three graphs) | int8 size: 135 MB
Usage
from optimum.onnxruntime import ORTModelForSeq2SeqLM
from transformers import AutoTokenizer
repo = "TigreGotico/opus-mt-tr-az-onnx"
tok = AutoTokenizer.from_pretrained(repo)
model = ORTModelForSeq2SeqLM.from_pretrained(repo, use_cache=True, use_merged=False) # fp32
# int8: ORTModelForSeq2SeqLM.from_pretrained(repo, subfolder="int8", use_cache=True, use_merged=False)
inputs = tok("Bugün hava çok güzel.", return_tensors="pt")
out = model.generate(**inputs, num_beams=4, max_new_tokens=64)
print(tok.decode(out[0], skip_special_tokens=True))
Parity with the original PyTorch model
10 general-domain Turkish sentences (source language for this tr->az
pair - the first attempt at this gate mistakenly used English sentences and
was discarded, since it measures nothing for a Turkish-source model),
exact-string-match of generated output against MarianMTModel.generate()
on the original Helsinki-NLP/opus-mt-tr-az checkpoint.
parity:
fp32_greedy: 1.00 # 10/10
fp32_beam4: 1.00 # 10/10
int8_greedy: 0.70 # 7/10
int8_beam4: 0.50 # 5/10
| Decoding | fp32 exact match | int8 exact match |
|---|---|---|
| greedy (num_beams=1) | 10/10 (100.0%) | 7/10 (70.0%) |
| beam=4 | 10/10 (100.0%) | 5/10 (50.0%) |
fp32 is a faithful reproduction of the original model at both decoding settings. int8 dynamic quantization causes some quality loss on this checkpoint. Prefer fp32; int8 is provided for size-constrained deployments where some quality loss is acceptable.
- Downloads last month
- 25
Model tree for TigreGotico/opus-mt-tr-az-onnx
Base model
Helsinki-NLP/opus-mt-tr-az