Telugu BERT Large (telugu-bert-large)

A BERT-style masked-language model pretrained from scratch on Telugu-only text. This card reports only measured numbers; anything that was not actually run is explicitly labeled NOT MEASURED rather than assumed or copied from other projects. See the "Honest evaluation" section below before using this model for any comparative or production claim.

Model Details

Property Value
Architecture BertForMaskedLM (standard BERT — absolute position embeddings; no RoPE/SwiGLU/pre-norm)
Parameters ≈268M (hidden 1024, 16 layers, 16 heads, intermediate 4096)
Vocabulary 64,000 WordPiece tokens, Telugu-trained tokenizer
Max sequence length 512
Precision float32 (model.safetensors, ~1.07 GB)
Training objective Masked language modeling (15% mask probability)
Framework 🤗 Transformers, transformers_version 4.36.0
Training hardware 8-GPU Kubernetes/PyTorchJob cluster (job telugu-bert-training-7jdg4)
Training length 3 epochs, 24,060 steps, effective batch size ≈2048, ~7h37m wall-clock

This is a Telugu-only encoder model (BertForMaskedLM). It is not a generative / chat model and was not trained on any other language.

Training Data

The model was pretrained on a locally assembled Telugu-only text corpus (~389 GB raw text on disk), combining:

  • Public web/curated corpora with Telugu-only filtering (large-scale Telugu web/news text, similar in spirit to Sangraha/OSCAR/mC4/CC-100 style Telugu subsets)
  • Telugu literature and traditional scripture text (e.g. Mahabharata/Ramayana/Puranas-style texts)
  • Telugu news and contemporary text

Note: this run consumed only 3 epochs over the corpus (~49.3M masked training examples seen), so it did not exhaust the full ~389 GB of available Telugu text — there is room to continue pretraining for a stronger model. No exact token count was independently re-verified for this specific card; do not treat "25-30B tokens" as a verified number for this checkpoint.

Honest Evaluation (measured, not aspirational)

All numbers below come from Kubernetes evaluation jobs run directly against this checkpoint (telugu-bert-large), documented in EVALUATION_REPORT.md in the training repo.

In-domain MLM diagnostic

⚠️ This is an in-domain diagnostic — the corpus sample was drawn from the same pretraining data distribution (no independent held-out split was reserved during pretraining), so treat this as an optimistic sanity check, not a generalization measurement.

Metric Value
Samples 4,096 sequences / 313,364 masked tokens
Loss 2.5262
Perplexity 12.51
Masked-token accuracy 54.30%

Downstream: Telugu NER (google/xtreme, config PAN-X.te)

A token-classification head was fine-tuned on the PAN-X.te 1,000-example train split (3 epochs) on top of the frozen-pretraining encoder, then evaluated on the untouched 1,000-example test split.

Metric telugu-bert-large (this model) kuppuluri/telugu_bertu (public baseline, same protocol)
Precision 66.06% 54.67%
Recall 73.78% 60.46%
Entity F1 69.71% 57.42%
Token accuracy 92.32% 88.09%
Eval loss 0.2742 0.3889

This is a genuine, reproducible head-to-head comparison run with the identical fine-tuning protocol on the same PAN-X.te split against a real public Telugu BERT baseline (kuppuluri/telugu_bertu), and telugu-bert-large scores higher on every reported NER metric in this specific comparison.

What was NOT measured

The following are not verified for this model and should not be assumed:

  • No held-out (leakage-free) MLM perplexity/accuracy — only the in-domain diagnostic above exists.
  • No sentiment analysis, topic classification, or question-answering benchmark was run (no labeled datasets or fine-tuning code for those tasks exist yet for this project).
  • No head-to-head comparison against BigBERT-Telugu, IndicBERT, mBERT, MuRIL, or XLM-R was performed (their weights were not available in this environment).
  • Any specific numeric target such as "MLM perplexity < 2.4712" or "sentiment accuracy > 91.94%" referenced in earlier drafts of this project's brief is an unverified target figure from the task prompt, not a measured result for this checkpoint or for any baseline actually tested here. It is intentionally omitted from this card.

Bottom line: this is credibly one of the largest from-scratch, Telugu-only pretrained BERT encoders publicly documented (≈268M params vs. ~89–110M for common public Telugu/Indic baselines), and it outperforms the kuppuluri/telugu_bertu public Telugu BERT on a real, identical-protocol PAN-X.te NER comparison. It should not be described as "state of the art" or as beating BigBERT/IndicBERT/mBERT overall, since those broader comparisons were never run.

Usage

Masked language modeling

from transformers import AutoTokenizer, AutoModelForMaskedLM
import torch

tokenizer = AutoTokenizer.from_pretrained("veeranool/telugu-bert-large")
model = AutoModelForMaskedLM.from_pretrained("veeranool/telugu-bert-large")

text = "తెలుగు [MASK] ద్రావిడ భాష."
inputs = tokenizer(text, return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

mask_idx = (inputs.input_ids == tokenizer.mask_token_id).nonzero(as_tuple=True)[1]
top5 = outputs.logits[0, mask_idx].topk(5, dim=-1).indices[0]
print([tokenizer.decode([t]) for t in top5])

Feature extraction / fine-tuning (e.g. NER, classification)

from transformers import AutoTokenizer, AutoModel

tokenizer = AutoTokenizer.from_pretrained("veeranool/telugu-bert-large")
model = AutoModel.from_pretrained("veeranool/telugu-bert-large")

inputs = tokenizer("హైదరాబాద్ తెలంగాణ రాష్ట్ర రాజధాని.", return_tensors="pt")
outputs = model(**inputs)
last_hidden_state = outputs.last_hidden_state  # (1, seq_len, 1024)

Use AutoModelForTokenClassification / AutoModelForSequenceClassification / AutoModelForQuestionAnswering with .from_pretrained("veeranool/telugu-bert-large") to attach a task head and fine-tune, exactly as was done for the PAN-X.te NER evaluation above.

Telugu-specific example (literature/scripture-style text)

text = "రామాయణం అనేది వాల్మీకి రచించిన ఒక [MASK] కావ్యం."

Limitations

  • Telugu only: trained exclusively on Telugu text; not suitable for other languages or code-mixed text without further adaptation.
  • Encoder-only (BERT) architecture: designed for masked-language-modeling and understanding/fine-tuning tasks (classification, NER, extractive QA), not open-ended text generation.
  • Undertrained relative to corpus size: only 3 epochs / 49M examples were consumed out of a much larger (389GB) available corpus; MLM loss had plateaued around 2.77 (training loss) at the end of the run, suggesting the model may benefit from a longer schedule.
  • No held-out generalization number for MLM: the only MLM metric available is an in-domain diagnostic (see above), not a clean validation/test split.
  • Only one downstream task evaluated: NER (PAN-X.te). Performance on sentiment analysis, topic classification, question answering, or other tasks is unknown.
  • Absolute position embeddings, 512 max length: like classic BERT, this model cannot process sequences longer than 512 tokens natively.

Ethical Considerations

  • The pretraining corpus is intended to be Telugu-only; large-scale web-scraped text may still contain noise, biases, or low-quality passages that were not manually audited.
  • Traditional/scripture-style Telugu text was included as part of the corpus; users applying the model to culturally or religiously sensitive content should apply their own review.
  • As with any BERT-style MLM, the model can reflect biases present in its training data. It has not been evaluated for bias/fairness properties.

Training Infrastructure

  • Pretraining: Kubernetes PyTorchJob, job name telugu-bert-training-7jdg4, 8 GPUs, bf16 mixed precision, per_device_train_batch_size=64, gradient_accumulation_steps=4, learning_rate=1e-4, num_train_epochs=3, mlm_probability=0.15.
  • Tokenizer: WordPiece, 64,000-token vocabulary, trained on the same Telugu corpus.
  • Evaluation: separate Kubernetes Jobs for the in-domain MLM diagnostic and the PAN-X.te NER fine-tuning/evaluation.

Citation

@misc{telugu-bert-large,
  title  = {Telugu BERT Large: A Telugu-only pretrained BERT encoder},
  author = {veeranool},
  year   = {2026},
  publisher = {Hugging Face},
  note   = {Evaluated in-domain MLM diagnostic and PAN-X.te NER fine-tuning; see model card for
            full, unembellished evaluation results.}
}
Downloads last month
27
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results