Instructions to use veeranool/telugu-bert-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use veeranool/telugu-bert-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="veeranool/telugu-bert-large")# Load model directly from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("veeranool/telugu-bert-large") model = AutoModelForMaskedLM.from_pretrained("veeranool/telugu-bert-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Telugu BERT Large (telugu-bert-large)
A BERT-style masked-language model pretrained from scratch on Telugu-only text. This card reports only measured numbers; anything that was not actually run is explicitly labeled NOT MEASURED rather than assumed or copied from other projects. See the "Honest evaluation" section below before using this model for any comparative or production claim.
Model Details
| Property | Value |
|---|---|
| Architecture | BertForMaskedLM (standard BERT — absolute position embeddings; no RoPE/SwiGLU/pre-norm) |
| Parameters | ≈268M (hidden 1024, 16 layers, 16 heads, intermediate 4096) |
| Vocabulary | 64,000 WordPiece tokens, Telugu-trained tokenizer |
| Max sequence length | 512 |
| Precision | float32 (model.safetensors, ~1.07 GB) |
| Training objective | Masked language modeling (15% mask probability) |
| Framework | 🤗 Transformers, transformers_version 4.36.0 |
| Training hardware | 8-GPU Kubernetes/PyTorchJob cluster (job telugu-bert-training-7jdg4) |
| Training length | 3 epochs, 24,060 steps, effective batch size ≈2048, ~7h37m wall-clock |
This is a Telugu-only encoder model (BertForMaskedLM). It is not a generative / chat model
and was not trained on any other language.
Training Data
The model was pretrained on a locally assembled Telugu-only text corpus (~389 GB raw text on disk), combining:
- Public web/curated corpora with Telugu-only filtering (large-scale Telugu web/news text, similar in spirit to Sangraha/OSCAR/mC4/CC-100 style Telugu subsets)
- Telugu literature and traditional scripture text (e.g. Mahabharata/Ramayana/Puranas-style texts)
- Telugu news and contemporary text
Note: this run consumed only 3 epochs over the corpus (~49.3M masked training examples seen), so it did not exhaust the full ~389 GB of available Telugu text — there is room to continue pretraining for a stronger model. No exact token count was independently re-verified for this specific card; do not treat "25-30B tokens" as a verified number for this checkpoint.
Honest Evaluation (measured, not aspirational)
All numbers below come from Kubernetes evaluation jobs run directly against this checkpoint
(telugu-bert-large), documented in EVALUATION_REPORT.md in the training repo.
In-domain MLM diagnostic
⚠️ This is an in-domain diagnostic — the corpus sample was drawn from the same pretraining data distribution (no independent held-out split was reserved during pretraining), so treat this as an optimistic sanity check, not a generalization measurement.
| Metric | Value |
|---|---|
| Samples | 4,096 sequences / 313,364 masked tokens |
| Loss | 2.5262 |
| Perplexity | 12.51 |
| Masked-token accuracy | 54.30% |
Downstream: Telugu NER (google/xtreme, config PAN-X.te)
A token-classification head was fine-tuned on the PAN-X.te 1,000-example train split (3 epochs) on top of the frozen-pretraining encoder, then evaluated on the untouched 1,000-example test split.
| Metric | telugu-bert-large (this model) |
kuppuluri/telugu_bertu (public baseline, same protocol) |
|---|---|---|
| Precision | 66.06% | 54.67% |
| Recall | 73.78% | 60.46% |
| Entity F1 | 69.71% | 57.42% |
| Token accuracy | 92.32% | 88.09% |
| Eval loss | 0.2742 | 0.3889 |
This is a genuine, reproducible head-to-head comparison run with the identical fine-tuning protocol
on the same PAN-X.te split against a real public Telugu BERT baseline
(kuppuluri/telugu_bertu), and telugu-bert-large scores higher on every reported NER metric in
this specific comparison.
What was NOT measured
The following are not verified for this model and should not be assumed:
- No held-out (leakage-free) MLM perplexity/accuracy — only the in-domain diagnostic above exists.
- No sentiment analysis, topic classification, or question-answering benchmark was run (no labeled datasets or fine-tuning code for those tasks exist yet for this project).
- No head-to-head comparison against BigBERT-Telugu, IndicBERT, mBERT, MuRIL, or XLM-R was performed (their weights were not available in this environment).
- Any specific numeric target such as "MLM perplexity < 2.4712" or "sentiment accuracy > 91.94%" referenced in earlier drafts of this project's brief is an unverified target figure from the task prompt, not a measured result for this checkpoint or for any baseline actually tested here. It is intentionally omitted from this card.
Bottom line: this is credibly one of the largest from-scratch, Telugu-only pretrained BERT
encoders publicly documented (≈268M params vs. ~89–110M for common public Telugu/Indic baselines),
and it outperforms the kuppuluri/telugu_bertu public Telugu BERT on a real, identical-protocol
PAN-X.te NER comparison. It should not be described as "state of the art" or as beating
BigBERT/IndicBERT/mBERT overall, since those broader comparisons were never run.
Usage
Masked language modeling
from transformers import AutoTokenizer, AutoModelForMaskedLM
import torch
tokenizer = AutoTokenizer.from_pretrained("veeranool/telugu-bert-large")
model = AutoModelForMaskedLM.from_pretrained("veeranool/telugu-bert-large")
text = "తెలుగు [MASK] ద్రావిడ భాష."
inputs = tokenizer(text, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
mask_idx = (inputs.input_ids == tokenizer.mask_token_id).nonzero(as_tuple=True)[1]
top5 = outputs.logits[0, mask_idx].topk(5, dim=-1).indices[0]
print([tokenizer.decode([t]) for t in top5])
Feature extraction / fine-tuning (e.g. NER, classification)
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("veeranool/telugu-bert-large")
model = AutoModel.from_pretrained("veeranool/telugu-bert-large")
inputs = tokenizer("హైదరాబాద్ తెలంగాణ రాష్ట్ర రాజధాని.", return_tensors="pt")
outputs = model(**inputs)
last_hidden_state = outputs.last_hidden_state # (1, seq_len, 1024)
Use AutoModelForTokenClassification / AutoModelForSequenceClassification /
AutoModelForQuestionAnswering with .from_pretrained("veeranool/telugu-bert-large") to attach a
task head and fine-tune, exactly as was done for the PAN-X.te NER evaluation above.
Telugu-specific example (literature/scripture-style text)
text = "రామాయణం అనేది వాల్మీకి రచించిన ఒక [MASK] కావ్యం."
Limitations
- Telugu only: trained exclusively on Telugu text; not suitable for other languages or code-mixed text without further adaptation.
- Encoder-only (BERT) architecture: designed for masked-language-modeling and understanding/fine-tuning tasks (classification, NER, extractive QA), not open-ended text generation.
- Undertrained relative to corpus size: only 3 epochs /
49M examples were consumed out of a much larger (389GB) available corpus; MLM loss had plateaued around 2.77 (training loss) at the end of the run, suggesting the model may benefit from a longer schedule. - No held-out generalization number for MLM: the only MLM metric available is an in-domain diagnostic (see above), not a clean validation/test split.
- Only one downstream task evaluated: NER (PAN-X.te). Performance on sentiment analysis, topic classification, question answering, or other tasks is unknown.
- Absolute position embeddings, 512 max length: like classic BERT, this model cannot process sequences longer than 512 tokens natively.
Ethical Considerations
- The pretraining corpus is intended to be Telugu-only; large-scale web-scraped text may still contain noise, biases, or low-quality passages that were not manually audited.
- Traditional/scripture-style Telugu text was included as part of the corpus; users applying the model to culturally or religiously sensitive content should apply their own review.
- As with any BERT-style MLM, the model can reflect biases present in its training data. It has not been evaluated for bias/fairness properties.
Training Infrastructure
- Pretraining: Kubernetes
PyTorchJob, job nametelugu-bert-training-7jdg4, 8 GPUs, bf16 mixed precision,per_device_train_batch_size=64,gradient_accumulation_steps=4,learning_rate=1e-4,num_train_epochs=3,mlm_probability=0.15. - Tokenizer: WordPiece, 64,000-token vocabulary, trained on the same Telugu corpus.
- Evaluation: separate Kubernetes Jobs for the in-domain MLM diagnostic and the PAN-X.te NER fine-tuning/evaluation.
Citation
@misc{telugu-bert-large,
title = {Telugu BERT Large: A Telugu-only pretrained BERT encoder},
author = {veeranool},
year = {2026},
publisher = {Hugging Face},
note = {Evaluated in-domain MLM diagnostic and PAN-X.te NER fine-tuning; see model card for
full, unembellished evaluation results.}
}
- Downloads last month
- 27
Evaluation results
- Entity F1 on PAN-X.tetest set self-reported0.697
- Precision on PAN-X.tetest set self-reported0.661
- Recall on PAN-X.tetest set self-reported0.738
- Token accuracy on PAN-X.tetest set self-reported0.923