AmeriXLM
Multilingual language model pre-trained on Latin American Spanish and Brazilian Portuguese, with additional coverage of English (Americas) and indigenous languages of the hemisphere.
Model Description
AmeriXLM corrects a systematic gap in existing multilingual models: XLM-R and mBERT allocate training compute proportional to available text volume, which underrepresents Latin American regional variants even within Spanish and Portuguese allocations. AmeriXLM is trained on a corpus that explicitly weights Latin American regional text.
- Architecture: RoBERTa
- Parameters: 270M
- Vocabulary: 64,000 tokens trained on regional corpus
- Context length: 512 tokens
Training Data
- Brazilian Portuguese: 45B tokens (news, government, legal, web)
- Latin American Spanish: 38B tokens (news, government, legal, web)
- English Americas: 12B tokens
- Indigenous languages: 800M tokens (Quechua, Guarani, Nahuatl)
Benchmark Results
| Benchmark | Score | XLM-R large baseline | Delta |
|---|---|---|---|
| AmeriNLI es-LA | 83.7 | 79.4 | +4.3 |
| AmeriNLI pt-BR | 82.9 | 78.2 | +4.7 |
| RegioNER es-LA F1 | 88.2 | 86.7 | +1.5 |
| RegioNER pt-BR F1 | 87.6 | 85.3 | +2.3 |
Usage
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("genia-americas/ameri-xlm") model = AutoModel.from_pretrained("genia-americas/ameri-xlm")
inputs = tokenizer("Texto en español latinoamericano", return_tensors="pt") outputs = model(**inputs)
Citation
@techreport{genia2025amerixlm, title={AmeriXLM: A Multilingual Language Model for the Americas}, author={GENIA Americas Corporation}, institution={GENIA Americas / RaceFor.AI}, year={2025}, url={https://github.com/GENIA-Americas/multimodal-ai-americas} }
Links
- Repository: https://github.com/GENIA-Americas/multimodal-ai-americas
- DOI: https://doi.org/10.5281/zenodo.20437260
- Platform: https://www.glapagos.com
- Network: https://www.racefor.ai