Part 1 V3
Training updates: 6300. Non-padding training tokens: 36302363.
V3 is a partial run and is not training-budget matched to the completed variants.
Custom decoder-only PyTorch model; no pretrained transformer. Shared 16,000-token byte-level BPE. Context capacity: 384. Translation prompts: <bos> <vi-or-ja> SOURCE <en>.
Install torch==2.7.0 and tokenizers==0.23.1. Download this repository, add its directory to sys.path, then use from load_model import load_model; model, tokenizer = load_model(directory). Call model.forward(input_ids, attention_mask); logits have shape [batch, time, 16000].
Training data: belumind/en-vi-ja-curated-500k-triplets, official training split. Checkpoint provenance and test evaluation settings are provided as JSON.
- Downloads last month
- 32