Model card: AnuLM-Translate-400M (ckpt_translate.pt)

An English ↔ Hindi sentence translator: the three-language AnuLM base (AnuLM-Base-400M, step 36,000) fine-tuned for one pass over two million sentence pairs. Full log and samples: docs/RESULTS.md §24 and docs/TRANSLATE_PLAN.md.

Licence: CC BY-NC 4.0, research use only. Samanantar is CC BY-NC 4.0 and the IIT Bombay corpus is released for research; the weights inherit those terms. Not affiliated with Sarvam AI, AI4Bharat, BharatGen or the Government of India.

Results

chrF on FLORES-200 devtest, all 1,012 sentences per direction, greedy:

direction this model untuned base NLLB-600M (reference)
English → Hindi 41.5 0.1 ~55
Hindi → English 43.4 3.1 ~55

Held-out answer loss fell 4.05 → 2.30 over the pass and was still improving at the last eval. Everyday sentences of 8 to 25 words translate well; rare names, numbers and long clauses are where it slips.

Data

source licence used
ai4bharat/samanantar, Hindi split CC BY-NC 4.0 most of the 2M pairs
cfilt/iitb-english-hindi research use the rest
FLORES-200 devtest CC BY-SA 4.0 evaluation only

Pairs were filtered for length, script and duplicates and used in both directions as English: … newline Hindi: … and the reverse, loss on the target sentence only. 4M items packed into 493,370 rows of 512; 61,672 steps at batch 8, learning rate 5e-5, about six hours on one RTX 5070 Ti.

Architecture and tokenizer

Same network as every 400M AnuLM checkpoint: 20 layers, hidden 1,024, GQA 16 query / 4 key-value heads, 24 routed experts of 192 with top-4 routing and aux-loss-free balancing, sliding window 256 on layers 0–9, context 512, 32k vocabulary. 397.7M parameters, 173.5M active per token. Tokenizer multi32k, a byte-level BPE trained on the Hindi + English + Python mix (docs/RESULTS.md §22).

Use

The page picks the direction from the script you type; the API takes mode: "translate". One sentence at a time; temperature 0.3 or lower. python serve.py --ckpt <this folder> then open http://127.0.0.1:8000. "Continue the text" and "answer a question" on this checkpoint are not useful: fine-tuning on translation pairs eroded the base model's free-text ability. Use AnuLM-Base-400M or AnuLM-Hindi-QA-400M for those.

How to load

With transformers

The modelling code travels with the weights, so trust_remote_code=True is required and there is nothing to clone:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("toonist/AnuLM-Translate-400M", trust_remote_code=True)
tok = AutoTokenizer.from_pretrained("toonist/AnuLM-Translate-400M")

ids = tok("English: Where is the nearest hospital?\nHindi:", return_tensors="pt")
print(tok.decode(model.generate(**ids, max_new_tokens=60, do_sample=False)[0]))

Greedy output is identical to the AnuLM repository's own generate, cached or not, and the tokenizer here agrees with bpe.py token for token; both are pinned by tests in that repository.

Load it in float32 — which the config now asks for, so the line above is enough. Do not force dtype=torch.bfloat16: this is a mixture-of-experts model whose router keeps a per-expert bias, and rounding that bias to 16 bits changes which experts fire. The output does not get slightly worse, it collapses into repeated tokens. The big tensors are stored in bfloat16 and upcast on load; the router bias and the norms are stored in float32 for this reason. For speed, use torch.autocast over float32 weights, which is how every number in this card was measured. The model loads the code from this repo at trust_remote_code=True; pin a revision if you want that fixed.

Batches must be unpadded (one sequence at a time), and beam search is not supported.

With the AnuLM repository

Every script there takes this folder wherever it takes a .pt:

hf download toonist/AnuLM-Translate-400M --local-dir AnuLM-Translate-400M
python serve.py  --ckpt AnuLM-Translate-400M          # web page at http://127.0.0.1:8000
python sample.py --ckpt AnuLM-Translate-400M --prompt "English: Where is the nearest hospital?\nHindi:"
from model import load_checkpoint, AnuLM
ck = load_checkpoint("AnuLM-Translate-400M")
m = AnuLM(ck["cfg"]).eval(); m.load_state_dict(ck["model"])

Files: model.safetensors (397.7M parameters), tokenizer.multi32k.json for bpe.BPE.load, and tokenizer.json in the tokenizers format for AutoTokenizer.

Downloads last month
434
Safetensors
Model size
0.4B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for toonist/AnuLM-Translate-400M

Finetuned
(3)
this model

Datasets used to train toonist/AnuLM-Translate-400M