AmeriXLM

Multilingual language model pre-trained on Latin American Spanish and Brazilian Portuguese, with additional coverage of English (Americas) and indigenous languages of the hemisphere.

Model Description

AmeriXLM corrects a systematic gap in existing multilingual models: XLM-R and mBERT allocate training compute proportional to available text volume, which underrepresents Latin American regional variants even within Spanish and Portuguese allocations. AmeriXLM is trained on a corpus that explicitly weights Latin American regional text.

  • Architecture: RoBERTa
  • Parameters: 270M
  • Vocabulary: 64,000 tokens trained on regional corpus
  • Context length: 512 tokens

Training Data

  • Brazilian Portuguese: 45B tokens (news, government, legal, web)
  • Latin American Spanish: 38B tokens (news, government, legal, web)
  • English Americas: 12B tokens
  • Indigenous languages: 800M tokens (Quechua, Guarani, Nahuatl)

Benchmark Results

Benchmark Score XLM-R large baseline Delta
AmeriNLI es-LA 83.7 79.4 +4.3
AmeriNLI pt-BR 82.9 78.2 +4.7
RegioNER es-LA F1 88.2 86.7 +1.5
RegioNER pt-BR F1 87.6 85.3 +2.3

Usage

from transformers import AutoTokenizer, AutoModel

tokenizer = AutoTokenizer.from_pretrained("genia-americas/ameri-xlm") model = AutoModel.from_pretrained("genia-americas/ameri-xlm")

inputs = tokenizer("Texto en español latinoamericano", return_tensors="pt") outputs = model(**inputs)

Citation

@techreport{genia2025amerixlm, title={AmeriXLM: A Multilingual Language Model for the Americas}, author={GENIA Americas Corporation}, institution={GENIA Americas / RaceFor.AI}, year={2025}, url={https://github.com/GENIA-Americas/multimodal-ai-americas} }

Links

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support