AnViMt
AnViMt — An=English Vi=Vietnamese Mt=Machine Translation — dịch EN → VI cho văn bản thường + kỹ thuật/hệ thống/lập trình/hội thoại.
Export sang ONNX để chạy chỉ với onnxruntime + sentencepiece + numpy (không cần torch/transformers).
Architecture
Base: transformers.BartForConditionalGeneration train từ đầu, không phải pretrained BART.
vocab_size=32000(SentencePiece BPE,byte_fallback=True)d_model=512|encoder_layers=6|decoder_layers=6encoder_ffn_dim=2048|decoder_ffn_dim=2048encoder_attention_heads=8|decoder_attention_heads=8max_position_embeddings=256|dropout 0.1tie_word_embeddings=True|is_encoder_decoder=Truepad_token_id=0(<pad>) |unk_token_id=1(<unk>) |bos_token_id=2(<s>) |eos_token_id=3(</s>) |decoder_start_token_id=2- Params: ~60.79M
EN text
↓ spm.model [2] + BPE ids + [3] (pad 0, bos 2, eos 3, max 256)
↓
encoder_model.onnx (input_ids, attention_mask → encoder_hidden_states [batch, seq, 512])
↓
decoder_model.onnx / decoder_with_past_model.onnx
(decoder_input_ids [2] + encoder_hidden_states + encoder_attention_mask + past_key_values → logits)
↓ autoregressive greedy (argmax) đến eos 3
↓ sp.decode → VI text
Files
Export bằng optimum task=seq2seq-lm opset=14:
AnViMt/
├── encoder_model.onnx # 136M
├── decoder_model.onnx # 223M
├── decoder_with_past_model.onnx # KV cache (tùy chọn, tăng tốc)
├── spm.model # 752K, vocab 32000
└── spm.vocab
Runtime tối thiểu chỉ cần encoder_model.onnx + decoder_model.onnx + spm.model. decoder_with_past_model.onnx dùng khi muốn KV-cache.
Không cần: config.json, pytorch_model.bin, safetensors, torch, transformers.
Installation
pip install onnxruntime sentencepiece numpy
# Termux: apt install python3-onnxruntime python3-sentencepiece python3-numpy
Quick Start
project/
├── main.py
├── encoder_model.onnx
├── decoder_model.onnx
└── spm.model # (decoder_with_past_model.onnx nếu có)
main.py thực tế (đã test, xử lý đúng BOS/EOS/pad):
import numpy as np, onnxruntime as ort, sentencepiece as spm
sp = spm.SentencePieceProcessor(); sp.load("spm.model")
enc = ort.InferenceSession("encoder_model.onnx", providers=["CPUExecutionProvider"])
dec = ort.InferenceSession("decoder_model.onnx", providers=["CPUExecutionProvider"])
BOS, EOS, PAD = 2, 3, 0
def translate(text, max_len=256):
ids = [BOS] + sp.encode(text, out_type=int)[:254] + [EOS]
input_ids = np.array([ids], dtype=np.int64)
attn = np.ones_like(input_ids)
enc_hs = enc.run(None, {"input_ids": input_ids, "attention_mask": attn})[0]
gen = [BOS]
for _ in range(max_len):
dec_ids = np.array([gen], dtype=np.int64)
logits = dec.run(None, {"input_ids": dec_ids, "encoder_hidden_states": enc_hs, "encoder_attention_mask": attn})[0]
nxt = int(np.argmax(logits[0, -1]))
gen.append(nxt)
if nxt == EOS: break
out = [x for x in gen[1:] if x not in (BOS, EOS, PAD)]
if EOS in out: out = out[:out.index(EOS)]
return sp.decode(out)
Training Data
Pretrain 1.13M EN-VI (PhoMT + TED + TECH) → finetune 200k TECH → finetune 27k v3 (hội thoại + kỹ thuật). Tokenizer BPE 32k train trên 400k pairs reservoir sampled, nmt_nfkc, split_digits, byte_fallback.
Limitations
EN → VI only. Câu >254 BPE tokens bị cắt. Thuật ngữ mới có thể giữ nguyên EN.
License
CC BY-NC 4.0 — non-commercial. Author: Phitran21