Part 1 V3

Training updates: 6300. Non-padding training tokens: 36302363.

V3 is a partial run and is not training-budget matched to the completed variants.

Custom decoder-only PyTorch model; no pretrained transformer. Shared 16,000-token byte-level BPE. Context capacity: 384. Translation prompts: <bos> <vi-or-ja> SOURCE <en>.

Install torch==2.7.0 and tokenizers==0.23.1. Download this repository, add its directory to sys.path, then use from load_model import load_model; model, tokenizer = load_model(directory). Call model.forward(input_ids, attention_mask); logits have shape [batch, time, 16000].

Training data: belumind/en-vi-ja-curated-500k-triplets, official training split. Checkpoint provenance and test evaluation settings are provided as JSON.

Downloads last month
32
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support